Papers with text transcriptions
Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech (2023.findings-acl)
Copied to clipboard
| Challenge: | Expressive text-to-speech aims to generate high-quality samples with rich prosody . prosodic attributes in highly dynamic voices are difficult to capture and model without intonation . |
| Approach: | They propose a pipeline that enhances prosody modeling and sampling by introducing a self-supervised masked autoencoder and a diffusion model to sample diverse prosodic patterns within the latent space. |
| Outcome: | The proposed pipeline achieves new state-of-the-art in text-to-speech with natural and expressive synthesis. |
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results . |
| Approach: | They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach . |
| Outcome: | The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts. |
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Speech Vecalign is a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. |
| Approach: | They propose a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. |
| Outcome: | The proposed method outperforms SpeechMatrix models on 3,000 hours of unlabeled speech documents and produces longer speech-to-speech alignments. |